Two Technologies Vie for Recognition in Speech Market
نویسنده
چکیده
I mprovements to voice-recognition algorithms and greater computing power have changed speech technology from an approach with limited uses to an increasingly important part of many applications. As this process has unfolded, developers have searched for an open standard , rather than proprietary development and runtime environments, that would let them easily and quickly add speech input/output capabilities to applications that function across platforms , said Inderpal Singh Mumick, founder and CEO of Kirusa, a wireless-platform developer. Two principal approaches have resulted from this search: VoiceXML and SALT (speech application language tags). Each approach's goal is to let users with minimal training add speech interfaces to applications. Both utilize speech-recognition and speech-synthesis systems to convert voice input to a format that a computer system can understand, to retrieve or produce the requested information, and to convert it back into speech for the user. VoiceXML is a completely self-contained language for writing speech-based user interfaces. SALT, a newer technology, is a collection of XML tags that developers can embed into an application written in other languages to create a speech or multimodal user interface. Multimodal interfaces can be controlled by voice and traditional input methods such as a keyboard or mouse. Both technologies let developers design speech applications that are portable across multiple platforms. " VoiceXML and SALT tags hide many of the low-level details often present in proprietary languages, " said Jim Larson, principal of Larson Technology Services and chair of the World Wide Web Consortium's (W3C's) Voice Browser Working Group. " This decreases the time required to write and debug code. And because standard languages are more widely used, there will be more developers trained in using them. " The speech technologies would be particularly useful for mobile applications because many handheld devices have tiny keyboards and screens, as well as other limitations that make traditional I/O approaches difficult, noted Rob Kassel, a product manager for SpeechWorks International and a representative of the SALT Forum, an industry consortium that promotes SALT use for multimodal and tele-phony applications. The stakes are huge. Allied Business Intelligence, an industry research firm, projects the global speech-technology market will increase from $677 million in 2002 to $897.8 million this year and to $5.3 billion by 2008, as Figure 1 shows. A number of forces are driving demand for speech technologies and, therefore, VoiceXML and SALT. Text-to-speech applications let users convert text-based data to speech, …
منابع مشابه
Persian Phone Recognition Using Acoustic Landmarks and Neural Network-based variability compensation methods
Speech recognition is a subfield of artificial intelligence that develops technologies to convert speech utterance into transcription. So far, various methods such as hidden Markov models and artificial neural networks have been used to develop speech recognition systems. In most of these systems, the speech signal frames are processed uniformly, while the information is not evenly distributed ...
متن کاملSpoken Term Detection for Persian News of Islamic Republic of Iran Broadcasting
Islamic Republic of Iran Broadcasting (IRIB) as one of the biggest broadcasting organizations, produces thousands of hours of media content daily. Accordingly, the IRIBchr('39')s archive is one of the richest archives in Iran containing a huge amount of multimedia data. Monitoring this massive volume of data, and brows and retrieval of this archive is one of the key issues for this broadcasting...
متن کاملImproving the performance of MFCC for Persian robust speech recognition
The Mel Frequency cepstral coefficients are the most widely used feature in speech recognition but they are very sensitive to noise. In this paper to achieve a satisfactorily performance in Automatic Speech Recognition (ASR) applications we introduce a noise robust new set of MFCC vector estimated through following steps. First, spectral mean normalization is a pre-processing which applies to t...
متن کاملImproved Bayesian Training for Context-Dependent Modeling in Continuous Persian Speech Recognition
Context-dependent modeling is a widely used technique for better phone modeling in continuous speech recognition. While different types of context-dependent models have been used, triphones have been known as the most effective ones. In this paper, a Maximum a Posteriori (MAP) estimation approach has been used to estimate the parameters of the untied triphone model set used in data-driven clust...
متن کاملComputer assisted foreign language teaching and learning by using speech technologies
Early works aiming at technical support of teaching and learning pronunciation of foreign language can be found in the 1980s, where technologies for speech analysis, synthesis, and recognition were examined. In the 1990s, with the help of statistical frameworks, speech recognition and synthesis technologies were greatly improved and CALL (Computer Assisted Language Learning) studies benefited m...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- IEEE Computer
دوره 36 شماره
صفحات -
تاریخ انتشار 2003